BackGPU OOM

GPU OOM

Xiaomi
2026-09-21 09:55:57

Xiaomi’s MiMo halts two public RL training runs after spending $3.47 million

Xiaomi has stopped both publicly streamed reinforcement learning training runs for its MiMo model, with Pro and Flash each completing 30 training steps as their last full run. Combined spending reached $3.4747 million, including $2.6207 million for Pro and $854,000 for Flash. The gains were notable: on DeepSWE v1.1, Pro rose from 58.41 at step 1 to 72.57, while Flash climbed from 48.67 to 65.68. Still, both curves saw clear pullbacks during training rather than moving up in a straight line. Other benchmark panels also improved, with Pro’s internal coding evaluation increasing from 57.54 to 65.43 and AutomationBench rising from 45.2 to 53.1. The process also ran into operational issues. Pro hit a GPU OOM caused by uneven expert load, and both the training cluster and evaluator experienced disconnects and restarts. Flash had to restart from step 15 because of an infrastructure error. Xiaomi also later filtered out tasks that had become too easy for Pro and removed the cyber dataset from later Pro training after abnormal rollout patterns appeared.

30
Xiaomi’s MiMo halts two public RL training runs after spending $3.47 million